Papers with data collection
The Margarita Dialogue Corpus: A Data Set for Time-Offset Interactions and Unstructured Dialogue Systems (2020.lrec-1)
Copied to clipboard
| Challenge: | Time-Offset Interaction Applications (TOIAs) simulate face-to-face conversations between humans and digital human avatars recorded in the past. |
| Approach: | They propose a methodology for creating the knowledge base for a TOIA, a dialogue corpus, and baselines for single-turn answer retrieval. |
| Outcome: | The proposed method lets the avatar maker list pairs by intuition, guessing what possible questions a user may ask to the avatar. |
Are Large Language Models Good at Lexical Semantics? A Case of Taxonomy Learning (2024.lrec-main)
Copied to clipboard
| Challenge: | Recent studies on LLMs do not pay enough attention to linguistic and lexical semantic tasks, such as taxonomy learning. |
| Approach: | They propose a method for stochastic graph traversal and a new algorithm for data collection . they propose LLaMA-2 and Mistral for a lexical semantic task . |
| Outcome: | The proposed models can perform linguistic and lexical tasks, but they lack basic skills in taxonomy learning. |
Analyzing Bayesian Crosslingual Transfer in Topic Models (N19-1)
Copied to clipboard
| Challenge: | a theoretical analysis of crosslingual transfer in probabilistic topic models is presented . we use Gibbs sampling to quantify the loss of knowledge across languages . |
| Approach: | They propose a method to quantify the loss of knowledge across languages during crosslingual transfer in probabilistic topic models. |
| Outcome: | The proposed model quantifies the loss of knowledge across languages during this process . it is validated on a diverse set of five languages and discusses best practices for data collection and model design . |
TweetTaglish: A Dataset for Investigating Tagalog-English Code-Switching (2022.lrec-1)
Copied to clipboard
| Challenge: | a large dataset is available to study Tagalog-English code-switching in low-resource settings. |
| Approach: | They propose to use a large dataset to investigate Tagalog-English code-switching . they use linguistic data from Tagalogue and Tagalit-English to investigate their results . |
| Outcome: | The proposed dataset achieves a strong performance benchmark for Tagalog-English code-switching. |
WarriorCoder: Learning from Expert Battles to Augment Code Large Language Models (2025.acl-long)
Copied to clipboard
Huawen Feng, Pu Zhao, Qingfeng Sun, Can Xu, Fangkai Yang, Lu Wang, Qianli Ma, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang, Qi Zhang
| Challenge: | Recent code large language models have demonstrated impressive performance on code-related tasks. |
| Approach: | They propose a paradigm that learns from expert battles to address these limitations . they create an arena where leading LLMs challenge each other with evaluations . |
| Outcome: | The proposed model improves on existing models by leveraging expert battles . it achieves state-of-the-art performance even without relying on proprietary models . |
Beyond Fair Pay: Ethical Implications of NLP Crowdsourcing (2021.naacl-main)
Copied to clipboard
| Challenge: | Ethical considerations regarding the use of crowdworkers are limited to labor conditions . the Final Rule did not anticipate the use online crowdsourcing platforms for data collection . |
| Approach: | They propose to reopen discussion regarding ethical use of crowdworkers in NLP research . they propose to use online crowdsourcing platforms to evaluate risk of harm . |
| Outcome: | The proposed study identifies common scenarios where crowdworkers performing NLP tasks are at risk of harm. |
UOUO: Uncontextualized Uncommon Objects for Measuring Knowledge Horizons of Vision Language Models (2024.emnlp-main)
Copied to clipboard
Xinyu Pi, Mingyuan Wu, Jize Jiang, Haozhen Zheng, Beitong Tian, ChengXiang Zhai, Klara Nahrstedt, Zhiting Hu
| Challenge: | Vision-Language Models (VLMs) perform on par with larger models in general domain visual grounding and question-answering benchmarks. |
| Approach: | They propose a "Uncontextualized Uncommon Objects" benchmark to evaluate their performance on common datasets. |
| Outcome: | The proposed benchmark focuses on systematically testing VLMs with both large and small parameter counts on rare and specialized objects. |
Don’t paraphrase, detect! Rapid and Effective Data Collection for Semantic Parsing (D19-1)
Copied to clipboard
| Challenge: | a major hurdle on the road to conversational interfaces is the difficulty in collecting data that maps language utterances to logical forms . crowdsourcing and crowdsourcing have been used to generate pseudo-language paired with logical form . however, this data collection method often leads to low performance on real data . |
| Approach: | They propose a method that uses crowdsourcing to map language utterances to logical forms . they quantify the effects of mismatches between the true and induced distributions . |
| Outcome: | The proposed method leads to 70.6 accuracy on the true distribution, compared to 51.3 in paraphrase-based data collection. |
LVLMs and Humans Ground Differently in Referential Communication (2026.acl-long)
Copied to clipboard
Peter Zeng, Weiling Li, Amie J. Paige, Zhengxiang Wang, Panagiotis Kaliosis, Dimitris Samaras, Gregory J. Zelinsky, Susan Brennan, Owen Rambow
| Challenge: | generative AI agents cannot model common ground in a way that enables smooth communication . a recent study examined whether large language models and large vision language models engage in grounding as human discourse partners do . |
| Approach: | They propose to use referential communication to model common ground between a pair of directors and a picture matching system. |
| Outcome: | The proposed experiment shows that generative AI agents cannot model common ground . human conversation relies on common ground accrued and updated by interacting partners . |
Does Putting a Linguist in the Loop Improve NLU Data Collection? (2021.findings-emnlp)
Copied to clipboard
Alicia Parrish, William Huang, Omar Agha, Soo-Hwan Lee, Nikita Nangia, Alexia Warstadt, Karmanya Aggarwal, Emily Allaway, Tal Linzen, Samuel R. Bowman
| Challenge: | Many datasets for training and evaluating natural language understanding (NLU) models contain systematic artifacts that are identified only after data collection is complete. |
| Approach: | They propose to have linguists identify artifacts and gaps in the data and communicate with non-expert crowdworkers to adjust task instructions and incentives. |
| Outcome: | The proposed protocol does not increase accuracy on out-of-domain test sets, and adds a chatroom does not. |
Matina: A Large-Scale 73B Token Persian Text Corpus (2025.naacl-long)
Copied to clipboard
Sara Bourbour Hosseinbeigi, Fatemeh Taherinezhad, Heshaam Faili, Hamed Baghbani, Fatemeh Nadi, Mostafa Amiri
| Challenge: | Existing Persian datasets are small and lack content diversity . lack of high-quality data has slowed development of NLP models and open-source LLMs for Persian. |
| Approach: | They propose a Persian dataset of 72.9B tokens that is preprocessed and deduplicated to ensure high data quality. |
| Outcome: | The proposed model performs well on key Persian NLP tasks. |
Towards Transferable Personality Representation Learning based on Triplet Comparisons and Its Applications (2025.emnlp-main)
Copied to clipboard
Kai Tang, Rui Wang, Renyu Zhu, Minmin Lin, Xiao Ding, Tangjie Lv, Changjie Fan, Runze Wu, Haobo Wang
| Challenge: | Existing methods for personality analysis treat corpus as a single unit for classification, but this approach presents several challenges. |
| Approach: | They propose a task paradigm for text-based personality representation learning that uses a triplet personality trend comparison dataset to learn single-sentence personality embeddings with desirable metric properties. |
| Outcome: | The proposed model significantly boosts performance across various applications, including personality detection, personality retrieval, and emotion translation prediction. |
Forecasting Future International Events: A Reliable Dataset for Text-Based Event Modeling (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing approaches for text-based event prediction are limited in quality due to dynamic nature of international relations and conflicting economic dynamics. |
| Approach: | They propose a novel dataset that leverages the advanced reasoning capabilities of large-language models to address these limitations. |
| Outcome: | The proposed dataset features high-quality scoring labels generated through advanced prompt modeling and rigorously validated by domain experts in political science. |
Researching Less-Resourced Languages – the DigiSami Corpus (L18-1)
Copied to clipboard
| Challenge: | DigiSami project aims to support research on endangered languages . it uses spoken corpus and speech technology for the Fenno-Ugric language North Sami . |
| Approach: | They describe the DigiSami project and its research results for the Fenno-Ugric language North Sami . they discuss ethical and privacy issues related to data collection for less-resourced languages and indigenous communities . |
| Outcome: | The DigiSami project focuses on spoken corpus collection and speech technology for the Fenno-Ugric language North Sami. |
Bootstrapping LLM-based Task-Oriented Dialogue Agents via Self-Talk (2024.findings-acl)
Copied to clipboard
| Challenge: | Large language models (LLMs) are powerful dialogue agents, but specializing them towards fulfilling a specific function can be prohibitive in terms of feasibility, time, and resources. |
| Approach: | They propose a method for training large language models by enabling "self-talk" they propose supervised fine-tuning of LLMs to improve quality of dialogues . |
| Outcome: | The proposed method generates training data via "self-talk" of LLMs that can be refined and utilized for supervised fine-tuning. |
ChatGLM-Math: Improving Math Problem-Solving in Large Language Models with a Self-Critique Pipeline (2024.findings-emnlp)
Copied to clipboard
Yifan Xu, Xiao Liu, Xinghan Liu, Zhenyu Hou, Yueyan Li, Xiaohan Zhang, Zihan Wang, Aohan Zeng, Zhengxiao Du, Zhao Wenyi, Jie Tang, Yuxiao Dong
| Challenge: | Large language models (LLMs) have shown excellent mastering of human language but struggle in real-world applications that require mathematical problem-solving. |
| Approach: | They propose a pipeline to train a general Math-Critique model from the LLM itself to provide feedback signals and employ rejective fine-tuning and direct preference optimization over the Llm's own generations for data collection. |
| Outcome: | The proposed pipeline outperforms existing LLMs that could be two times larger. |
Mitigating Translationese in Low-resource Languages: The Storyboard Approach (2024.lrec-main)
Copied to clipboard
Garry Kuwanto, Eno-Abasi E. Urua, Priscilla Amondi Amuok, Shamsuddeen Hassan Muhammad, Anuoluwapo Aremu, Verrah Otiende, Loice Emma Nanyanga, Teresiah W. Nyoike, Aniefon D. Akpan, Nsima Ab Udouboh, Idongesit Udeme Archibong, Idara Effiong Moses, Ifeoluwatayo A. Ige, Benjamin Ajibade, Olumide Benjamin Awokoya, Idris Abdulmumin, Saminu Mohammad Aliyu, Ruqayya Nasir Iro, Ibrahim Said Ahmad, Deontae Smith, Praise-EL Michaels, David Ifeoluwa Adelani, Derry Tanti Wijaya, Anietie Andy
| Challenge: | Low-resource languages often face challenges in acquiring high-quality language data due to the reliance on translation-based methods, which introduce the translationese effect. |
| Approach: | They propose a method that uses storyboards to elicit more fluent and natural sentences from native speakers without direct exposure to the source text. |
| Outcome: | The proposed method compared with traditional translation-based methods in terms of accuracy and fluency. |
Dial HEALTHDIAL for Advice: A Multilingual and Multi-Parallel Spoken Dialogue Dataset for Knowledge-Grounded Information Seeking (2026.findings-acl)
Copied to clipboard
Songbo Hu, Yinhong Liu, Ej Zhou, Evgeniia Razumovskaia, Xiaobin Wang, Alexander Fraser, Ivan Vulić, Anna Korhonen
| Challenge: | Creating spoken dialogue datasets is methodologically challenging due to the personally identifiable nature of speech signals. |
| Approach: | They propose a large-scale, multilingual, and multi-parallel dataset for developing and evaluating retrieval-augmented generation-based spoken dialogue systems. |
| Outcome: | The proposed dataset includes 6,000 information-seeking dialogues and 163 hours of user speech recorded from native speakers of four official WHO languages. |